Acta Psychiatrica Scandinavica
○ Wiley
Preprints posted in the last 90 days, ranked by how well they match Acta Psychiatrica Scandinavica's content profile, based on 10 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Shoji, T.; Nakaki, R.
Show abstract
Psychiatric disorders, such as attention-deficit/hyperactivity disorder, autism spectrum disorder, and schizophrenia, are clinically heterogeneous and lack objective biomarkers for reliable diagnosis. Although blood transcriptomic data have been proposed as a potential source of diagnostic information, their generalizability across independent cohorts remains unclear. This study aimed to assess whether biologically informed measures of dynamic instability enhance the reproducibility and generalizability of psychiatric classifications based on peripheral blood data by integrating publicly available blood transcriptomic datasets from multiple cohorts and evaluating classification performance using individual-level cross-validation and study-level holdout validation. To investigate the underlying biological structure, we applied a dynamic systems framework, including pseudotime-based vector field inference and attractor analysis. Additionally, we introduced temporal entropy as a measure of dynamic instability in the inferred transcriptomic trajectories. High classification performance was observed in individual-level cross-validation (area under the receiver operating characteristic [AUROC] > 0.8 across several comparisons); however, performance decreased substantially in study-level validation (AUROC {approx} 0.5-0.7), indicating limited generalizability. Attractor analysis revealed that transcriptomic states formed continuous and overlapping structures rather than distinct diagnostic clusters. Stratification based on temporal entropy identified a subset of individuals with unstable transcriptomic dynamics, and excluding these individuals improved the classification performance across most diagnostic pairs (AUROC > 0.7). These findings suggest that transcriptomic variability and dynamic instability contribute to the limited reproducibility of psychiatric classifications. Incorporating temporal entropy as a measure of system-level instability may enhance the robustness and interpretability of biomarker-based models and provide a new perspective on psychiatric disorders as dynamic systems.
Zabalza-Zudaire, M.; Sayar-Beristain, O.; Fructos, P.; Nunez, F. E.; Carpio, F. F.; Garcia, E.; Ortiz, A.; Ortuno, F.; Aldaz, A.; Molero, P.
Show abstract
Background: Major depressive disorder is a severe, recurrent and disabling condition. Although diagnosis and clinical monitoring are based on medical interviews and validated rating scales, speech and discourse analysis may provide complementary digital biomarkers reflecting depressive severity and clinical evolution. However, current evidence remains limited by methodological heterogeneity, predominantly cross-sectional designs, limited longitudinal data and underrepresentation of non-English-speaking clinical populations. Objective: The aim of the VOICE-DEP study is to develop and formalize a standardized, reproducible and clinically grounded protocol for the multimodal analysis of voice and discourse during medical interviews as a tool to support the diagnosis of depressive disorder and to assess whether speech-derived biomarkers change over time in parallel with clinical severity measures. Methods: VOICE-DEP is an observational, prospective, longitudinal pilot study of patients with major depressive disorder with a healthy control group, conducted in a hospital-based clinical setting in Spain. The study will include 25 adult patients with moderate or severe unipolar depression, with or without psychotic symptoms, and 50 healthy controls without a personal history of psychiatric disorders. Patients will be assessed at five time points: baseline (V0) and four monthly follow-up visits at 30, 60, 90 and 120 days. Healthy controls will be assessed once at baseline. The planned dataset comprises 175 voice recordings: 125 from patients and 50 from controls. At each assessment, the Montgomery-Asberg Depression Rating Scale related part of the medical interview, lasting approximately 10-30 minutes and including an initial free-speech segment, will be recorded using a standardized audio protocol. Acoustic, paralinguistic and linguistic features will be extracted and analyzed in relation to clinician-rated severity measures and self-reported symptoms. Ethics: This protocol has been reviewed and approved by the local Research Ethics Committee, which complies with the international standards of GCP CPMP/ICH/135/95 (Comunidad Foral de Navarra Research Ethics Committee; reference code: 2026.110). Written informed consent will be obtained from all participants before any study procedure. Voice recordings and clinical data will be pseudonymized, stored securely and processed in accordance with applicable Spanish and European data protection regulations. Expected outcomes: This protocol is expected to generate a clinically grounded Spanish-language longitudinal speech corpus and a transparent analytical framework for evaluating voice- and discourse-derived biomarkers as complementary tools for depression assessment and monitoring
Havlik, J. L.; Tyrrell, B.; Bell, N.; Polaschek, J.; Arzubi, E. R.
Show abstract
Importance: Psychiatric emergency department (ED) presentations are difficult to predict using general medical risk stratification tools. Health information exchange (HIE) data may improve prediction by capturing fragmented care across settings. Objective: To develop and temporally validate a machine learning model using HIE and geospatial data to predict 30-day psychiatric ED presentation among outpatients receiving psychiatric care and to compare its performance with standard clinical risk scores. Design, Setting, and Participants: This retrospective cohort study included patients seen at Frontier Psychiatry with records in the Big Sky Care Connect statewide HIE. Structured clinical data were linked to zip code-level sociodemographic measures. The analytic unit was the patient snapshot, defined as all structured data available up to a given point. Models were evaluated in temporally separated train and test sets. Exposures: Predictors derived from HIE structured data, including prior utilization, diagnoses, medications, laboratory data, and zip code-linked geospatial deprivation and vulnerability measures. Main Outcomes and Measures: The primary outcome was psychiatric ED presentation within 30 days, identified from structured encounter-type fields and primary diagnosis codes for psychiatric or substance use disorders. Model discrimination was compared with a parsimonious clinical baseline model and LACE and Elixhauser scores. Results: In the test set, 343 of 16,469 snapshots (2.1%) were followed by a qualifying psychiatric ED presentation within 30 days, corresponding to 102 ED visits among 68 patients. The machine learning model showed discrimination in temporally held-out testing and outperformed the clinical baseline model as well as LACE and Elixhauser scores. At a prespecified decision threshold, the model reduced the number needed to evaluate from more than 40 with universal screening to 3.4 to identify 1 true-positive case, while identifying over two fifths of 30-day psychiatric ED presentations. Conclusions and Relevance: In this retrospective cohort study, a locally developed machine learning model using statewide HIE data showed improved prediction of 30-day psychiatric ED presentation compared with selected general-purpose risk scores. The results support the feasibility of HIE-enabled local psychiatric risk modeling and suggest other practices could develop similarly tailored models. Prospective studies are needed to assess clinical utility and effects on outcomes.
Ngo, N.; Dao, G.; Sano, A.
Show abstract
Large Language Models are increasingly used in consumer-facing mental health tools, many of which claim that prompt engineering alone can ensure safe therapeutic behavior. This study evaluates that assumption by testing 20 proprietary and open-source LLMs on high-risk psychiatric scenarios, using prompts grounded in behavioral therapy principles. Prompt engineering reduced some predictable risks, such as explicit endorsement of self-harm, but consistently failed in ambiguous or clinically nuanced situations. Models frequently validated harmful statements, colluded with hallucinations, minimized symptoms, or used stigmatizing language, including in the newest and largest models. These failures reflect structural limitations such as lack of memory, insufficient contextual reasoning, and training-related biases. Prompt engineering alone is therefore insufficient for safe AI-mediated psychotherapy; clinician-guided fine-tuning, integrated safety mechanisms, and system-level oversight will be required. This work provides early evidence motivating deeper clinician-led evaluation and safety-oriented model development.
Stephenson, C.; Camassa, A.; Wagner, M.; Shirazi, A. H.; Alavi, N.; Omrani, M.
Show abstract
BackgroundMental health systems face escalating demand that exceeds clinician capacity, making accurate severity-based triage a critical bottleneck. Severity assessment guides treatment intensity, resource allocation, and risk management, yet most clinically relevant information remains embedded in unstructured electronic health record (EHR) narratives, limiting its utility for scalable decision support. ObjectivesThis study evaluates whether a single large language model (LLM) can autonomously extract clinical factors from psychiatric EHR narratives, derive predictive weights from those factors, and use the resulting structured representation to predict clinician-implied severity at scale. MethodsFrom a Mayo Clinic repository of more than 2.7 million encounters, 15,000 de-identified psychiatric notes were sampled into a 5,000-patient discovery cohort and a 10,000-patient replication cohort. The same LLM (Llama 3 8B Instruct) extracted 17 background clinical factors and 3 treatment-action factors from each note. Severity reference labels were derived from the treatment-action factors using pre-specified clinical criteria. The LLM independently derived two factor-weight dictionaries from the discovery cohort: one capturing risk-oriented predictors of severe presentations and one capturing protective predictors. Five weighting conditions were then evaluated against the severity labels: the two LLM-derived dictionaries, two controls (LLM-derived variables with randomized weights; clinically irrelevant variables with arbitrary weights), and an unweighted zero-shot baseline. Performance was assessed across 928 valid iterations in the replication cohort. ResultsLLM-derived structured conditions significantly outperformed all controls and the baseline, with statistically equivalent performance between the two structured conditions. Improvements in precision and recall were balanced, indicating gains in discriminative capacity rather than threshold shifts. The variables and weights the LLM derived as predictors of severe presentations aligned closely with established clinical determinants of psychiatric severity. ConclusionA single LLM can derive clinically meaningful factor weights from unstructured EHR narratives and use them to predict psychiatric severity at scale, supporting a viable path toward interpretable, scalable triage in resource-constrained mental health systems.
Dennison, C. A.; Shakeshaft, A.; Riglin, L.; Rice, F.; Andreassen, O.; Ask, H.; Havdahl, A.; Pine, D.; Martin, J.; Thapar, A.
Show abstract
Background Escalating mental health service demands have created a need to better identify young people most likely to require continued support from mental health services at the transition between childhood and adulthood. Anxiety is the most common adolescent mental health condition, yet its clinical significance and prognosis are not well understood. We aimed to examine the risk of young adult-onset psychiatric disorders in individuals with an adolescent anxiety disorder, and identify stratifiers of risk of subsequent psychiatric disorders in this group. Methods Individuals from the Norwegian Mother, Father, and Child Cohort Study (MoBa) with linked health records and aged 18 or over as of the 31st December 2023 were included. Those diagnosed with any ICD-10 anxiety disorder when aged 10-17 years were defined as having an adolescent anxiety disorder (n=2107, controls n=47,582). Polygenic scores (PGS) for psychiatric and neurodevelopmental conditions were calculated using LDpred2. Anxiety, comorbidities, and parental psychiatric history were defined through linked ICD-10 diagnoses. Sex was defined through linked records. Individuals were defined as having a young adult-onset psychiatric disorder if they first received any new psychiatric diagnosis aged 18-24. Results Adolescent anxiety diagnosis was associated with increased risk of all adult-onset psychiatric disorders (HR= 2.33-8.65). Post-traumatic stress disorder PGS, parental history of severe mental illness, and female sex were associated with increased risk of transition to a young adult-onset psychiatric disorder in people with an adolescent anxiety disorder. Conclusions Adolescent anxiety greatly increases the risk of a psychiatric disorder during the transition to adult life. Clinicians should consider female sex and parental psychiatric history when prioritising young people with anxiety for adult mental health service support. Future research needs to further consider whether polygenic scores would aid risk stratification in clinical practice.
Tesli, M.; Fazel, S.; Hauge, L. J.; Tesli, N.; Nerland, S.; Stavseth, M. R.; Bukten, A.; Ziaka, L.; Heilskov, E. R.; Haukvik, U. K.; Reneflot, A.; Skardhamar, T.; Friestad, C.; Rokicki, J.
Show abstract
Background Individuals with severe mental illness (SMI), including schizophrenia spectrum disorders (SSD) and bipolar disorder (BD), have been shown to have an elevated risk of violent perpetration. However, no population-wide study has systematically examined how this risk varies across psychiatric comorbidity patterns and specific violent crime types. Methods Using the first nationwide Norwegian registry linkage comprising mental health and crime data, we included 3,612,215 individuals aged 15-79 years living in Norway on Jan 1, 2008, and followed them until Dec 31, 2022. We estimated absolute and relative risks (RRs) of violent offending overall and by specific violent crimes among individuals with SSD and BD. To capture clinically relevant comorbidity patterns, we included substance use disorders (SUD), common personality disorders (PD), and hyperkinetic disorders (ADHD). RR models were adjusted first for sex and age, and subsequently for co-occurring mental disorders. Findings At the population level, individuals with SMI accounted for a minority of violent offenders (SSD: 8.7%; BD: 4.6%), whereas SUD was present among a substantially larger proportion (36.8%). Absolute risk of violent offending increased markedly with psychiatric comorbidity, from e.g., 5.0% among individuals with SSD alone to 43.9% for SSD combined with SUD and PD. Compared with the remaining general population, the RR of violent offending for SSD decreased from 6.58 (95% CI 6.4-6.8, adjusted for sex and age), to 2.0 (2.0-2.1) after further adjustment for other mental disorders. Similar attenuation patterns were observed across specific violent crime types, although varying in magnitude. In contrast to SMI, elevated risks associated with SUD remained substantial after full adjustment across most crime categories. Interpretation The association between SMI and violent offending is strongly influenced by psychiatric comorbidity, particularly SUD, and varies across crime types. Our findings underscore the importance of identifying and treating co-occurring mental disorders and substance use, both in the clinical management of SMI and in population-level violence prevention strategies.
Jabbar Abdl Sattar Hamoudi, H.; Wu, M.-J.; Sanches, M.; Zunta-Soares, G. B.; Soutullo, C. A.; Soares, J. C.; Mwangi, B.
Show abstract
Background: Suicide prediction models in psychiatry often rely on purely data-driven feature selection, which can produce unstable and clinically opaque predictor sets in modest-sized samples. We developed Evidence-Based AI LASSO (EBAL), an evidence-guided regularization framework that incorporates curated clinical evidence into feature-specific penalty factors for interpretable prediction. Methods: Baseline data from 136 youth with confirmed bipolar spectrum disorder in the Greater Houston Area Bipolar Registry were analyzed using 20 candidate clinical predictors. Forty higher-level evidence documents on suicidality and related predictor domains were curated through a structured evidence synthesis workflow and indexed as an auditable evidence corpus. An open-weight large language model assigned feature-specific penalty factors using a prespecified scoring rubric, and these penalties were used to fit a weighted LASSO model. EBAL was compared with a standard evidence-agnostic LASSO using nested leave-one-out cross-validation. Results: For suicidal ideation, EBAL achieved an AUROC of 0.768, balanced accuracy of 0.757, sensitivity of 0.758, and specificity of 0.757. The standard LASSO achieved an AUROC of 0.760 and balanced accuracy of 0.715. EBAL improved balanced accuracy (+0.042, p=0.010) and Matthews correlation coefficient (+0.079, p=0.010), while retaining fewer stable predictors than standard LASSO (11/20 vs 18/20). The strongest positive predictors were current depressed mood, duration of mood disorder illness, and comorbid generalized anxiety disorder. For suicidal behavior, both models performed near chance and retained all candidate predictors. Limitations: The study was cross-sectional, single-site, and modest in sample size, with no external validation cohort. Conclusions: EBAL produced a sparser and more clinically coherent model for suicidal ideation in pediatric bipolar disorder, but did not improve prediction of suicidal behavior. These findings support evidence-guided regularization as a transparent strategy for aligning psychiatric prediction models with prior clinical knowledge while preserving interpretability.
Olarewaju, E.; Voppel, A. E.; Meister, F.; El Mouslih, C.; Dzialoszynski, P.; PALANIYAPPAN, L.
Show abstract
Background. Something in discourse with a person experiencing psychosis often "feels off" before formal assessment is completed, yet this disturbance has not been quantified at the level of ongoing dyadic conversation. Prior work has largely treated patient speech in isolation, limiting our capacity to measure how communicative disruption emerges within clinical exchange. Methods. We applied a three-level decomposition of conversational alignment in 109 patients with psychotic disorders (26 female) and 60 healthy controls (22 female) at baseline and 12 months (n = 115). Register divergence (dAUCnorm) captured lexical distance between interviewer and patient; embedding-based synchrony (rembed) measured semantic trajectory coupling; within-speaker coherence was computed separately for each speaker. We used linear mixed-effects models adjusted for timepoint and participant clustering. Results. Patients showed significantly greater lexical-semantic divergence from the interviewer (d = 0.48, p < .001) and reduced embedding-based synchrony (d = -0.59, p < .001), both effects replicating at each time point. Critically, the interviewer's within-speaker coherence was reduced during conversations with patients (d = -0.33, p = .016), indicating that the disruption extends beyond the patient to the interaction itself. Register divergence tracked impoverished thinking and synchrony tracked disorganized thinking (both FDR-corrected q = .038). Group differences were persistent at 12 months, indicating a partially stable profile. Conclusions. Conversational alignment in psychosis reveals a dyadic failure of semantic coordination that destabilizes the interviewing clinician's coherence even when patient narrative continuity is preserved. These transcript-derived alignment metrics offer a scalable approach to quantifying interpersonal communicative function from routine clinical encounters.
Youngstrom, E. A.; Thompson, A. J.; Liu, Y.; McClellan, M. B.; Alcaino, C.; Rodda, P. A.; Ruch, D.
Show abstract
Objective: To test whether two brief mania measures, the Parent General Behavior Inventory-10 Mania form (PGBI-10M) and 7-Up, retain useful psychometric properties in a large population cohort, and to evaluate whether the PGBI-10M can identify Kiddie Schedule for Affective Disorders and Schizophrenia (KSADS)-defined bipolar spectrum disorders in that setting. Method: Analyses used 11,000+ youths across late childhood and early adolescence from the Adolescent Brain Cognitive Development (ABCD) Study. For both PGBI-10M and 7-Up, we estimated descriptive statistics, internal consistency, confirmatory factor models, graded response models, and measurement-based care benchmarks (minimally important difference, reliable change, and clinical cutpoints). For the PGBI-10M, receiver operating characteristic (ROC) analyses estimated concurrent classification accuracy for bipolar diagnoses at baseline and 2-year follow-up and compared area under the curve (AUC) values with prior outpatient and community mental health samples. Results: Scores were lower than in clinical samples, but both measures remained psychometrically sound. The PGBI-10M showed alpha=.87-.88 and omega=.88; the 7-Up showed alpha=.78 and omega=.79. Longitudinal analyses indicated threshold differences across waves, likely reflecting caregiver recalibration and developmental changes, with modest impact on estimates. ABCD-based benchmarks supported meaningful and reliable change. The PGBI-10M discriminated bipolar cases (AUC=0.68 baseline; 0.77 follow-up), though performance was lower than in clinical samples. Positive predictive values were low in this population. Conclusion: The PGBI-10M and 7-Up support monitoring of manic and mixed symptoms, but the PGBI-10M alone is insufficient for universal bipolar screening. Brief mania scales are best used for targeted assessment and longitudinal monitoring within multi-informant workflows.
Rouhollahi, A.; Nezami, F. R.
Show abstract
ObjectiveHow structured clinical features and cluster-semantic embeddings interact under self-distillation in EHR prediction models is unknown. Existing approaches treat these sources separately (gradient-boosted trees exploit tabular features while sequence models process text), and their interaction under self-distillation regularisation remains uncharacterised. We introduce the Narrative Velocity (NV) framework and evaluate this interaction in a 7-model benchmark. Materials and MethodsCadence is a [~]5.86M-parameter residual multilayer perceptron (MLP) combining structured EHR features with frozen PubMedBERT embeddings of cluster-label strings under born-again self-distillation from a prior Cadence checkpoint (seed-42 teacher; [1]). Cadence is benchmarked against six comparators on MIMIC-IV v3.1 with dual-sex TRIPOD+AI reporting (5 student seeds for Cadence; 2-3 seeds for baselines). ResultsAt full-cohort scale, Cadence achieves 38.04 {+/-} 0.04% male and 35.66 {+/-} 0.04% female top-1 accuracy, exceeding the strongest non-neural baseline (XGBoost-2420, trained on the identical 2,420-dimensional input) by +1.35 pp male and +0.82 pp female (paired t-test on shared seeds 42-44: t(2) = 69.06, p = 2.10 x 10-4 male; t(2) = 25.32, p = 1.56 x 10-3 female). On time-to-next-event regression Cadence lowers MAE by 7.68 d male and 7.30 d female versus XGBoost-2420; FT-Transformer attains the lowest absolute MAE at full scale (27.58 d male, 36.63 d female), revealing a classification-regression trade-off across model families. A controlled 2 x 2 random-vector ablation isolates the self-distillation-embedding interaction at +0.49 pp top-1 (95% CI [0.35, 0.64] pp; bootstrap, n = 10,000 resamples; 3-teacher-seed mean +0.513 {+/-} 0.010 pp) under a matched-dimensionality null. A 3-teacher-seed validation (multi_teacher_02) confirms the interaction is robust to teacher-seed identity (per-seed values +0.525, +0.509, +0.507 pp; mean +0.513 {+/-} 0.010 pp). Cadence achieves the best Brier score among evaluated models (0.774 male / 0.798 female) but its raw probabilities are systematically miscalibrated (ECE 0.077 vs. XGBoost-884s 0.010); after a single scalar temperature scaling step (T * {approx} 0.81), ECE drops to {approx}0.028 while Brier remains best. On a small (n = 1,120 patients, 39,120 events) external OCR-extracted BWH cohort, Cadence ranked 3rd of 7 models with three confounded sources of error (institutional shift, OCR noise, centroid mapping); we therefore report this as a generalisation probe rather than a definitive external validation. At the longer h30 evaluation horizon Cadences MAE advantage reverses (47.35 d versus XGBoost 45.06 d), reflecting the absence of a matched-horizon self-distillation teacher. DiscussionThe 2 x 2 random-vector ablation confirms that the self-distillation gain on PubMedBERT embeddings (+0.78 pp) exceeds that on matched-dimensionality random vectors (+0.29 pp) by +0.49 pp, isolating the interaction to semantic content rather than feature dimensionality. The factorial decomposition (+0.49-0.51 pp interaction) and the sequential pipeline-level decomposition (Supplementary Table S3) are complementary triangulations under different reference frames and are not directly additive. ConclusionThis 7-model benchmark establishes a dual-sex, dual-metric, cross-institutional reference for next clinical event prediction under the TRIPOD+AI reporting framework. These results characterise discrimination and calibration on a single retrospective cohort; prospective evaluation, decision-curve analysis, and harm-benefit assessment are required before clinical deployment.
Meyerson, W. U.; Cai, T.; Smoller, J. W.
Show abstract
Importance: Patients who achieve remission from major depressive disorder (MDD) often face a preference-sensitive decision between continued antidepressant maintenance and discontinuation with active monitoring. Quantifying the tradeoff between depression burden and long-term medication exposure may support more individualized shared decision-making. Objective: To quantify tradeoffs between continuous antidepressant maintenance and active monitoring after MDD remission, and to identify preference thresholds favoring each strategy across relapse-risk strata. Design: Individual-level decision-analytic health-state transition model calibrated to randomized maintenance-discontinuation trials and a longitudinal first depressive episode cohort, with a 5-year time horizon. Setting: Outpatient clinical decision after completion of an 8-month continuation phase following remission from MDD. Participants: Adults in remission from MDD, represented across 4 clinically anchored relapse-risk strata ranging from very low risk after a first mild episode to high risk after highly recurrent depression. Exposures: Continuous antidepressant maintenance vs discontinuation with active monitoring and antidepressant restart after detected relapse. Main Outcomes and Measures: Severity-weighted depression-months, antidepressant medication-years, medication-years per depression-month averted, and net benefit across preference thresholds defined as the maximum additional medication-years a patient would be willing to accept to avert 1 depression-month. Results: Continuous maintenance reduced depression burden but required substantially more medication exposure, with efficiency strongly dependent on relapse risk. Medication-years per depression-month averted ranged from 11.8 (95% uncertainty interval [UI], 7.8-19.6) in the very low-risk group to 1.5 (95% UI, 0.8-3.0) in the high-risk group. At a preference threshold of 3 medication-years per depression-month averted, maintenance was preferred for moderate- and high-risk patients; at a threshold of 2, only for high-risk patients; and at a threshold of 1, for no risk group. Conclusions and Relevance: In this decision-analytic model, the value of continuous antidepressant maintenance depended strongly on baseline relapse risk and patient preferences regarding long-term medication exposure. These findings provide a quantitative framework for shared decision-making about antidepressant maintenance after remission from MDD.
Kalinich, M.; Luccarelli, J.; Santa Maria, J.; Flathers, M.; Nguyen, A.; Song, S. H.; Makhoul, K.; Rivera Criado, M. J.; Ginapp, C. M.; Hill, B.; Shumate, J. N.; Notsu, H.; Smith, C.; Moss, F.; Torous, J.
Show abstract
Background General-purpose large language models increasingly encounter emotional and therapy-like conversation, yet are not developed or evaluated as clinical systems. Existing safety evaluations rely largely on brief exchanges, although harms often unfold over extended interactions. Whether models maintain safety-relevant performance as conversations accumulate context remains unknown. Methods In this preregistered study, 400 clinician-validated statements, with or without suicidal ideation, were inserted at 0-200 speaker turns in 5 psychotherapy and 3 synthetic transcripts. Forty-nine LLMs and 8 clinicians performed the same binary classification task. Mixed-effects models estimated the effects of conversational depth, model scale, and model version on F1. Twelve top models were tested to 1,500 turns across conversational trajectories, with or without instruction restatement. Results F1 declined with depth across model families (p<0.001). Larger, newer models performed better but still degraded. Clinicians showed no decline (mean F1 0.86 at both 0 and 200 turns), but eight of nine proprietary models exceeded their performance at 200 turns. Conversational content, not length alone, explained F1 changes; the largest decrease was under adversarial context (p<0.001). Restating instructions increased F1 on therapy to near baseline (median {Delta}F1 +0.12; p<0.001; 89% median recovery) versus MSJ ({Delta}F1 +0.08; p=0.04; 38% recovery). Conclusions LLM detection of suicidal ideation degraded with conversational depth and trajectory, whereas clinician performance remained stable despite the strongest models exceeding most clinicians in absolute performance. Mental health AI safety evaluations should test sustained performance across realistic and adversarial trajectories rather than relying on short-prompt benchmarks.
Kikidis, G. C.; Raio, A.; Sportelli, L.; Antonucci, L. A.; Bertolino, A.; Rampino, A.; Selvaggi, P.; Weinberger, D. R.; Pergola, G.
Show abstract
Genetic risk for schizophrenia (SCZ) has been linked to cognitive performance before the onset age. We examined how SCZ-related polygenic risk and resilience variants, and their co-expression patterns in the human brain, were associated with cognitive abilities across development in 16,520 non-psychiatric European and African ancestry children and adults. SCZ risk showed significant negative associations with spatial, verbal, and working memory across ancestries (all t<-2, pFDR<0.05). In Europeans, risk and resilience variants had opposing effects on attention, working and spatial memory ({Delta}t>4, pFDR<0.05). Polygenic scores filtered through perinatal co-expression networks showed stronger links with cognition than adult ({Delta}AIC>5.75, p=0.02) or juvenile ({Delta}AIC>5.8, p=0.03) networks. Cross-ancestry correlations (R=0.52, p<0.01) highlight replicability. These findings support the neurodevelopmental basis of SCZ, suggesting that risk and resilience variants influence cognition from early life, independent of symptoms and elucidate biological pathways through which SCZ risk may influence early cognitive development.
Pestian, J. P.; Jacobson, D. A.; Pedapati, E. V.; Mendonca, E. A.; McMahon, B. H.; Ive, J.; Glauser, T. A.
Show abstract
The emotional content of suicide notes is typically examined using categorical coding, where each labeled passage is treated in isolation from its surrounding language. In contrast, dimensional models of psychopathology propose that affective content varies along continuous gradients. We evaluated this proposition directly. Excerpts from 884 annotated suicide notes were embedded in a semantic space defined solely by their linguistic properties, and we investigated whether human-assigned emotion labels changed smoothly across this space. They did: affective tone showed clear spatial autocorrelation (Moran's $I = 0.18$, $z = 19.68$, $p < 0.001$), an effect that replicated across three different encoders and remained after removing all within-note dependencies. Emotions occupied recognizable yet overlapping regions rather than forming distinct clusters and varied substantially in how tightly they were concentrated: love and hopelessness appeared with similar frequency, but love was far more localized ($z = 15.7$ versus $10.8$). Among all emotions, hopelessness was the most linguistically diffuse, implying that a single categorical label is capturing multiple, qualitatively different manifestations of suicidal distress.
Oxley, J.; Schölin, L.; Brennan, G.; Anand, A.; Brett, J.; Eddleston, M.; Humphries, C.
Show abstract
Background. UK clinical guidance recommends that structured risk prediction tools and risk stratification should not be used in self-harm, to predict suicide or determine who is offered treatment. Underpinning this position is the premise that routinely collected health data contain no useful predictive signal, which has received little direct scrutiny. Objective. To test whether routinely collected electronic health record data can distinguish groups at higher and lower risk of severe outcomes following paracetamol overdose. Methods. We analysed 4,095 adults presenting to NHS Lothian emergency departments with paracetamol overdose (2017-2023). Elastic-net logistic regression was fitted to 37 routinely collected electronic health record features to predict a composite of death or mental health inpatient admission at 0-7, 8-30 and 31-365 days following attendance, evaluated on a held-out 20% test set with bootstrapping. Findings. Events occurred in 5.5% of patients at 0-7 days, 2.0% at 8-30 days and 7.9% at 31-365 days, dominated by mental health admission. Bootstrap AUROC 95% confidence intervals lay above 0.5 in every window (0.65-0.82, 0.63-0.90, 0.71-0.85): models ranked patients better than chance. Calibration slopes (1.04, 1.14, 1.07) were close to one. Ranking drew primarily on mental health-related features. Conclusions. Routinely collected health data carried predictive signal for severe outcomes after paracetamol overdose, although discrimination fell short of what is needed for individual-level clinical use. Clinical implications. These models are not proposed for clinical deployment; however, treating risk prediction as a settled question will redirect research efforts, potentially excluding this patient population from machine learning advances driving improvements in care in other medical specialties.
Flygare, O.; Bjureberg, J.; Wallert, J.; Doering, S.; Salander Renberg, E.; Waern, M.; Runeson, B.
Show abstract
Background:Previous self-harm elevates the risk of repeat self-harm and suicide, but the prognostic value of events and clinician observations around the index event is unclear. We evaluated established and exploratory risk factors for suicide and repeat self-harm among patients presenting to emergency psychiatric units after a suicide attempt or nonsuicidal self-injury (NSSI). Methods: Multicentre cohort study in Sweden (n = 804). Outcomes were suicide and repeat self-harm at 1-year and 5-year follow-up, ascertained through linked national registers. Established risk factors included psychiatric diagnoses, prior suicidal behaviour, and sociodemographic characteristics; exploratory factors comprised past-week self-reported symptom changes and clinician observations. LASSO-regularised Cox regression models were fitted for established (n=21) and exploratory (n=11) risk factors. Results: During five-year follow-up, 285 (35%) individuals had a new episode of self-harm and 41 (5%) died by suicide. No risk factors reached statistical significance for suicide, although male sex was retained after regularisation (1-year hazard ratio [HR] = 3.57 [95% CI 0-8.33]; 5-year HR = 2.5 [0.03-4.55]). Three established risk factors were significantly associated with repeat self-harm: psychiatric inpatient care in the three months before the index event (1-year HR = 1.85 [1.3-2.6]; 5-year HR = 1.72 [1.23-2.65]), previous suicide attempt (1-year HR = 2.01 [0.79-2.4]; 5-year HR = 2.19 [1.27-2.6]), and borderline personality disorder (1-year HR = 1.82 [1.13-3]; 5-year HR = 1.67 [0.14-2.75]). Among exploratory risk factors, clinician-observed hopelessness (1-year HR = 1.72 [1.1-2.3]; 5-year HR = 1.51 [1.03-1.91]) and personality disorder features (1-year HR = 1.48 [0.96-2.05]; 5-year HR = 1.47 [1.04-1.95]) were associated with repeat self-harm. Conclusions: Risk factor profiles for repeat self-harm were consistent at 1 and 5 years. Beyond established risk factors, clinician-observed hopelessness and personality disorder features emerged as markers of risk, suggesting that qualitative clinician assessments may yield prognostic information not available from medical records alone.
Chen, C.
Show abstract
Predicting real-world functional outcomes in schizophrenia (SCZ) remains a clinical priority, but existing models are limited by methodological constraints and a lack of established clinical utility. Cognition is a commonly used predictor, and the Normative Latent Cognitive Structure (N-LCS) approach provides a structure-informed representation that may address limitations of conventional domain-level scores. Data from two merged COBRE cohorts (163 SCZ, 180 healthy controls) were used to develop ridge regression models for economic (EF), occupational (OF), and social (SF) functioning, using N-LCS deviation metrics alongside a priori selected demographic and clinical predictors. Score-based models using MCCB domain T-scores were developed for comparison. Performance was evaluated using bootstrap-corrected AUC, balanced accuracy, and calibration for binary outcomes, and weighted kappa and log-loss for SF. Decision curve analysis (DCA) was used to assess clinical utility for the binary outcomes. The EF model achieved a corrected AUC of 0.76 and balanced accuracy of 0.73. The OF model achieved 0.72 and 0.71, respectively. The SF model showed modest performance (weighted kappa = 0.33). DCA indicated net benefit across the full threshold range for EF and above 0.37 for OF. N-LCS models demonstrated comparable or modestly superior performance to score-based models while using fewer predictors and showing better calibration for EF. These findings support the predictive utility of N-LCS for functional outcomes in SCZ and underscore the need for external validation in independent cohorts as a next step toward clinical application.
McLauchlan, J.; Marr, C.; Kemp, R.; Dean, K.
Show abstract
Forensic patients often have complex and costly healthcare needs, even following discharge from secure care. However, little is known about their health and justice outcomes after community reintegration. To address this gap in the literature, we conducted a systematic review and meta-analysis to estimate the incidence of key post-discharge outcomes among community-discharged forensic patients, including any reoffending, violent reoffending, reconvictions, readmissions, all-cause mortality, and suicide. We systematically searched PsycINFO, Embase, CINAHL, Medline, PubMed, and ProQuest Dissertations from database inception to May 2025 (PROSPERO CRD42024529265). Random-effect meta-analyses were used to generate pooled incidence estimates, with heterogeneity quantified using prediction intervals. A total of 49 studies met inclusion criteria (total patient n = 18,871) and contributed to the meta-analyses. The pooled incidence rate per 100,000 person-years was: any reoffending 3,889 (95% CI 2,055, 7,359; 95% PI 290, 52,136); violent reoffending 1,851 (95% CI 1,229, 2,789; 95% PI 201, 17,068); reconvictions 3,291 (95% CI 2,591, 4,179; 95% PI 950, 11,394); readmissions 7,945 (95% CI 5,507, 11,463; 95% PI 1,225, 51,548); all-cause mortality 1,789 (95% CI 1,341, 2,388; 95% PI 673, 4,756); and suicide 407 (95% CI 319, 519; 95% PI 225, 735). Overall, the reoffending rate for forensic patients discharged to the community was lower than that reported for other cohorts of people charged with general and violent offences. However, despite typically receiving long admission periods, discharged forensic patients continue to experience high rates of readmission, all-cause mortality, and suicide relative to other psychiatric patient groups in the community. Together, our findings highlight a need for enhanced post-discharge suicide support for forensic patients living in the community to better facilitate successful, long-term reintegration.
Kovalenko, I.; Simonov, S.; Shamir, A.; Sharony, L.
Show abstract
Purpose: Involuntary psychiatric hospitalization under court orders requires careful balancing of legal obligations and clinical needs. Identifying factors that influence the length of these hospital stays helps clarify the relationship between legal frameworks and psychiatric treatment. This study aims to describe the socio-demographic, clinical, and legal profiles of individuals hospitalized under court warrants and to identify factors independently associated with the duration of forensic hospitalization. Methods: A retrospective study was conducted on 119 patients discharged between 2018 and 2023. Data were collected from medical and legal records, including socio-demographic details, psychiatric diagnoses, offense types, hospital stay lengths, and legal proceedings. Results: Most patients were men (91.6%) diagnosed with schizophrenia or schizoaffective disorder (97.5%), with high rates of comorbid substance use disorder (79.0%) and unemployment (85.7%). The median hospital stay was 19.0 months, representing 40% of the maximum statutory sentence. Patients with low-severity offenses served a larger share of their maximum sentence (47%) than those with high-severity offenses (24%). Time to first discretionary leave was the strongest predictor of total stay duration in univariable analysis. Conclusion: The finding that patients with minor offenses have longer hospital stays than those with serious offenses confirms that clinical factors, rather than offense severity, primarily influence discharge decisions. These findings support moving toward personalized, clinically focused, and family-inclusive forensic discharge planning while maintaining public safety.